Repository navigation
SS-455 - source: commit Kafka offsets on change and on a timer - #39462
Conversation
QA LLM Review1. MEDIUM -- Idle partitions are never re-committed, so the broker expires their consumer-group offsets
With the DetailsOn Suggested fix: keep the dedup, but re-commit unchanged offsets once the last successful commit is older than a refresh interval set well below typical broker retention (for example one hour, or a dyncfg). Because #39461 switched to |
7e97474 to
e1893ed
Compare
Confirmed and addressed |
QA LLM Review1. MEDIUM --
|
patrickwwbutler
left a comment
There was a problem hiding this comment.
Small question, plus suggestion to fix the LLM lint, but no blockers! LGTM
| # Low default so CI exercises the periodic recommit, which production | ||
| # only reaches after ten minutes. | ||
| VariableSystemParameter( | ||
| "kafka_offset_commit_refresh_interval", "10s", ["1s", "10s", "10m"] |
There was a problem hiding this comment.
| "kafka_offset_commit_refresh_interval", "10s", ["1s", "10s", "10m"] | |
| "kafka_offset_commit_refresh_interval", "10s", ["1s", "10s", "600s"] |
| // Zero disables the refresh. `interval` panics on a zero period, so the tick arm | ||
| // below is what honors it. | ||
| let mut refresh = | ||
| tokio::time::interval(refresh_interval.max(Duration::from_millis(1))); |
There was a problem hiding this comment.
if the value was set to 0, and this clamps it to 1ms, is that sustainable? I suppose it would just happen every iteration of the loop, which is okay?
There was a problem hiding this comment.
1ms would just update as fast as the calls would allow it (because the missed tick behavior is delay).
If someone sets it to 0 to disable, then the branch never fires because the precondition would return false:
materialize/src/storage/src/source/kafka.rs
Lines 732 to 735 in 97b4ca9
https://docs.rs/tokio/latest/tokio/macro.select.html
Evaluate all provided expressions. If the precondition returns false, disable the branch for the remainder of the current call to select!. Re-entering select! due to a loop clears the “disabled” state.
e1893ed to
97b4ca9
Compare
Brokers expire a standalone group's offset per partition offsets.retention.minutes after its last commit, so offsets the dedupe skips must still be recommitted periodically. The interval is the kafka_offset_commit_refresh_interval dyncfg, 10 minutes by default and 10 seconds in CI, with zero disabling the refresh. Co-Authored-By: Claude <noreply@anthropic.com>
97b4ca9 to
c220506
Compare
Each worker commits only when its offsets change, plus on a timer. The timer is needed because the consumer only ever assigns partitions, which results in the brokers treating the group as standalone. For standalone groups, brokers expire a partition's offset offsets.retention.minutes after its last commit, whether or not it is still current. The interval is the kafka_offset_commit_refresh_interval dyncfg: 10 minutes by default, 10 seconds in CI, zero disables it. A failed commit is retried on the next resume upper.
🤖 Generated with Claude Code
👨 Improved by human